Papers with hypothesis generation
Reflection in the Dark: Exposing and Escaping the Black Box in Reflective Prompt Optimization (2026.acl-srw)
Copied to clipboard
| Challenge: | Automatic prompt optimization (APO) is a powerful paradigm for improving LLM performance without manual prompt engineering. |
| Approach: | They propose a framework that decouples hypothesis generation from prompt rewriting . they propose VISTA framework that recovers accuracy to 87.57% on same defective seed . |
| Outcome: | The proposed framework outperforms baselines on GSM8K and AIME2025 on a defective seed. |
EVIDENCEMINER: Textual Evidence Discovery for Life Sciences (2020.acl-demos)
Copied to clipboard
Xuan Wang, Yingjun Guan, Weili Liu, Aabhas Chauhan, Enyi Jiang, Qi Li, David Liem, Dibakar Sigdel, John Caufield, Peipei Ping, Jiawei Han
| Challenge: | EVIDENCEMINER is a web-based system that allows users to query a natural language statement and retrieve textual evidence from a background corpora for life sciences. |
| Approach: | They propose a web-based system that lets users query a natural language statement and automatically retrieves textual evidence from a background corpora for life sciences. |
| Outcome: | EVIDENCEMINER is a web-based system that lets users query a natural language statement and automatically retrieves textual evidence from a background corpora for life sciences. |
Literature Meets Data: A Synergistic Approach to Hypothesis Generation (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods for hypothesis generation are theory-driven and data-driven, but they lack the computational power to complement each other. |
| Approach: | They develop a method that combines literature-based insights with data to perform LLM-powered hypothesis generation. |
| Outcome: | The proposed method outperforms baseline methods on five datasets and shows human accuracy improves on deception detection and AI generated content detection tasks. |
LMdiff: A Visual Diff Tool to Compare Language Models (2021.emnlp-demo)
Copied to clipboard
| Challenge: | LMdiff visually compares probability distributions of two different language models . notably absent from the range of available tools are those that aim to compare distributions produced by different models. |
| Approach: | They propose a tool that visually compares probability distributions of two different language models that differ through finetuning, distillation, or simply training with different parameter sizes. |
| Outcome: | The proposed tool allows the generation of hypotheses about model behavior by investigating text instances token by token and further assists in choosing interesting text instances from large corpora. |
DIXITWORLD: Evaluating Multimodal Abductive Reasoning in Vision-Language Models with Multi-Agent Dixit Gameplay (2026.acl-short)
Copied to clipboard
Yunxiang MO, Tianshi Zheng, Qing Zong, Jiayu Liu, Baixuan Xu, Yauwai Yim, Chunkit Chan, Jiaxin Bai, Yangqiu Song
| Challenge: | Existing evaluations of multimodal abductive reasoning are limited to static, single-agent tasks. |
| Approach: | They propose a multiagent evaluation suite that deconstructs the current evaluations of multimodal abductive reasoning in vision–language models. |
| Outcome: | The evaluation suite is based on two core components: DixitArena and DixitsBench. |
Advancing Abductive Reasoning in Knowledge Graphs through Complex Logical Hypothesis Generation (2024.acl-long)
Copied to clipboard
| Challenge: | Abductive reasoning is the process of making educated guesses to provide explanations for observations. |
| Approach: | They propose a task of complex logical hypothesis generation to generate a complex logique hypothesis that can explain a set of observations. |
| Outcome: | The proposed model generates logical hypotheses closer to the reference hypothesis, but not better on unseen observations. |
EA2E: Improving Consistency with Event Awareness for Document-Level Argument Extraction (2022.findings-naacl)
Copied to clipboard
| Challenge: | Recent work on document-level event argument extraction models each individual event in isolation and therefore causes inconsistency among extracted arguments across events. |
| Approach: | They propose an event-aware argument extraction model with augmented context to improve consistency . they hypothesize that participants tend to play consistent roles across multiple events in a document . |
| Outcome: | The proposed model improves consistency and accuracy of arguments extracted from documents. |
On the Role of Model Prior in Real-World Inductive Reasoning (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies have evaluated the inductive reasoning capabilities of Large Language Models (LLMs) by evaluating their ability to generate textual hypotheses based on in-context input-output pairs and test these hypothese based upon unseen examples. |
| Approach: | They evaluated three inductive reasoning strategies across five real-world tasks with three LLMs and found that hypothesis generation is primarily driven by the model’s inherent priors. |
| Outcome: | The proposed models generate high-quality hypotheses that can generalize to new instances when guided by in-context demonstrations. |
EvoNarrator: Modeling Scientific Evolution for Feasible Hypothesis Generation (2026.acl-long)
Copied to clipboard
| Challenge: | Scientific discovery evolution does not occur ex nihilo but is characterized by structural deepening and reconfiguration of existing functionalities. |
| Approach: | They propose a framework for hypothesis generation based on evolutionary narratives . they extract structured P-M-L-F quadruples from citation networks and introduce a mechanism to assess their semantic compatibility. |
| Outcome: | The proposed framework reduces logical disconnects by evaluating its semantic compatibility. |
From Hypothesis to Publication: A Comprehensive Survey of AI-Driven Research Support Systems (2025.findings-emnlp)
Copied to clipboard
Zekun Zhou, Xiaocheng Feng, Lei Huang, Xiachong Feng, Ziyun Song, Ruihan Chen, Liang Zhao, Weitao Ma, Yuxuan Gu, Baoxin Wang, Dayong Wu, Guoping Hu, Ting Liu, Bing Qin
| Challenge: | rapid development of artificial intelligence (AI) technologies has inspired researchers to explore how AI can accelerate and enhance research. |
| Approach: | They organize the relevant studies into three main categories: hypothesis formulation, hypothesis validation, and manuscript publication. |
| Outcome: | The authors summarize the current state of research in three main areas: hypothesis formulation, hypothesis validation, and manuscript publication. |
Context-Aware Reasoning On Parametric Knowledge for Inferring Causal Variables (2025.findings-emnlp)
Copied to clipboard
| Challenge: | randomized experiments provide strong inferences, but are often infeasible due to ethical or practical constraints. |
| Approach: | They propose a benchmark where the objective is to complete a partial causal graph. |
| Outcome: | The proposed benchmarks show that they can hypothesize backdoor variables between a cause and its effect. |
Many Heads Are Better Than One: Improved Scientific Idea Generation by A LLM-Based Multi-Agent System (2025.acl-long)
Copied to clipboard
Haoyang Su, Renqi Chen, Shixiang Tang, Zhenfei Yin, Xinzhe Zheng, Jinzhe Li, Biqing Qi, Qi Wu, Hui Li, Wanli Ouyang, Philip Torr, Bowen Zhou, Nanqing Dong
| Challenge: | Recent AI methods have shown promise in tasks such as hypothesis generation and experimental design, but they fail to replicate the collaborative nature of real-world scientific practices. |
| Approach: | They propose a virtual scientific system that mimics the collaborative nature of scientific research by organizing a team of agents to generate, evaluate, and refine research ideas. |
| Outcome: | The proposed system outperforms the state-of-the-art method in producing new scientific ideas and offers valuable insights to guide future research. |
OPINE: A Prior-calibrated Scoring Framework for LLM-based Multi-label Scientific Opinion Classification (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for scientific opinion classification rely on direct label generation and are limited by the multi-label nature of scientific expressions. |
| Approach: | They propose a framework that reformulates scientific opinion classification as a controllable pipeline. |
| Outcome: | The proposed framework outperforms baseline models on 18 discourse functions in micro, macro, and example settings. |
Debate-of-Thoughts: Resolving Knowledge Conflicts in LLMs Through Internal Deliberation (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods for retrieval augmented generation are based on a simplistic binary choice of relying on external contexts or memory. |
| Approach: | They propose a framework that transforms conflict resolution into an active deliberation process by incorporating contradictions as opportunities for deeper reasoning. |
| Outcome: | Experiments show that DoT outperforms state-of-the-art methods while generating transparent debate transcripts that explain its decisions. |
Teaching Language Models to Forecast Research Success Through Comparative Idea Evaluation (2026.findings-acl)
Copied to clipboard
| Challenge: | Language models are accelerating scientific research by automating hypothesis generation and implementation. |
| Approach: | They ask whether LMs can forecast the empirical success of research ideas before experiments . they frame evaluation as a reasoning task via Reinforcement Learning with Verifiable Rewards . |
| Outcome: | The proposed model outperforms off-the-shelf models in 77.1% of the evaluations . the model outpersforms GPT-5 in the evaluation of 11,488 idea pairs . |
GUIDE: Towards Scalable Advising for Research Ideas (2026.acl-long)
Copied to clipboard
| Challenge: | Existing systems that provide detailed, constructive feedback on academic papers struggle with review fidelity. |
| Approach: | They explore factors that underlie the development of robust advising systems . large language models have shown remarkable progress in tasks from text generation to code synthesis . |
| Outcome: | The proposed model outperforms general-purpose language models in acceptance rates for self-ranked top-30% submissions to ICLR 2025. |